nvshmem4py - #674
Draft
jeffhammond wants to merge 12 commits into
Draft
nvshmem4py#674jeffhammond wants to merge 12 commits into
jeffhammond wants to merge 12 commits into
Conversation
…d-to-end cuda-python/cuda-core promoted its Device/system API out of the cuda.core.experimental namespace into stable cuda.core sometime after these scripts were written, and renamed system.num_devices (an attribute) to system.get_num_devices() (a method) along the way. Both PYTHON/hello-nvshmem.py and PYTHON/nstream-cupy-nvshmem.py were still importing from the old experimental namespace, so neither could even import against a current install. Set up a local .venv (PRK/.venv, gitignored) with nvshmem4py-cu13 + cuda-python + cupy-cuda12x + mpi4py -- nvshmem4py itself isn't on the system Python and isn't packaged in this repo's own dependency tooling, so a project-local venv keeps it isolated. Verified both scripts run correctly under `mpirun -n 2` with this venv: hello-nvshmem.py completes and prints OK on both PEs, and nstream-cupy-nvshmem.py runs the full STREAM triad and validates. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The prior commit's "hack to disable jit?" (409c868) added a nvshmem.barrier() after every loop iteration because timings looked wrong without it. Investigated with a real, working nvshmem4py install (see previous commit): the actual bug is that CuPy operations issued with no explicit stream run on CuPy's own default stream, a different stream than the cuda.core `stream` object used for NVSHMEM's barrier/sync -- so stream.sync() was never actually waiting for the triad kernels to finish, and the CPU-time timer (see below) raced far ahead of the real GPU work. The per-iteration barrier "fixed" the symptom only by forcing the host to block/spin on every iteration, which happened to make time.process_time() advance too -- at the cost of also timing NVSHMEM barrier overhead every iteration instead of just the compute. Also switched the timer from time.process_time() (CPU time) to timeit.default_timer() (wall clock): process_time is the right choice for the CPU-bound sibling scripts (nstream.py, nstream-numpy.py, etc., where process/wall time track closely), but is fundamentally wrong for this GPU-async one -- it barely advances while the host blocks on the device. Fixed by binding CuPy's current stream to the same cuda.core stream NVSHMEM uses (cupy.cuda.Stream.from_external(stream), which cuda.core's Stream supports natively via the __cuda_stream__ protocol -- no deprecated ExternalStream needed), removing the per-iteration barrier, and switching to a wall-clock timer. Verified: `mpirun -n 2 python3 PYTHON/nstream-cupy-nvshmem.py 10 16777216` and a larger `20 134217728` run both validate correctly and report bandwidth in the range of real H100 HBM bandwidth (944 GB/s and 2.7 TB/s respectively, scaling up sensibly as problem size grows past launch-overhead-dominated territory) -- versus ~16 GB/s before this fix, which was dominated by per-iteration barrier overhead rather than measuring the actual triad. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
WIP